<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Voice computing</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Voice_computing"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Voice_computing rootpage-Voice_computing skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Voice computing</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<p><b>Voice computing</b> is the discipline that develops hardware or software to process voice inputs.<sup id="cite_ref-1" class="reference"><a href="#cite_note-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p><p>It spans many other fields including <a href="Human-computer_interaction" class="mw-redirect" title="Human-computer interaction">human-computer interaction</a>, <a href="Conversational_computing" class="mw-redirect" title="Conversational computing">conversational computing</a>, <a href="Linguistics" title="Linguistics">linguistics</a>, <a href="Natural_language_processing" title="Natural language processing">natural language processing</a>, <a href="Automatic_speech_recognition" class="mw-redirect" title="Automatic speech recognition">automatic speech recognition</a>, <a href="Speech_synthesis" title="Speech synthesis">speech synthesis</a>, <a href="Audio_engineering" class="mw-redirect" title="Audio engineering">audio engineering</a>, <a href="Digital_signal_processing" title="Digital signal processing">digital signal processing</a>, <a href="Cloud_computing" title="Cloud computing">cloud computing</a>, <a href="Data_science" title="Data science">data science</a>, <a href="Ethics" title="Ethics">ethics</a>, <a href="Law" title="Law">law</a>, and <a href="Information_security" title="Information security">information security</a>.
</p><p>Voice computing has become increasingly significant in modern times, especially with the advent of <a href="Smart_speakers" class="mw-redirect" title="Smart speakers">smart speakers</a> like the <a href="Amazon_Echo" title="Amazon Echo">Amazon Echo</a> and <a href="Google_Assistant" title="Google Assistant">Google Assistant</a>, a shift towards <a href="Serverless_computing" title="Serverless computing">serverless computing</a>, and improved accuracy of <a href="Speech_recognition" title="Speech recognition">speech recognition</a> and <a href="Text-to-speech" class="mw-redirect" title="Text-to-speech">text-to-speech</a> models.
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="History">History</h2></div>
<p>Voice computing has a rich history.<sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> First, scientists like <a href="Wolfgang_Kempelen" class="mw-redirect" title="Wolfgang Kempelen">Wolfgang Kempelen</a> started to build speech machines to produce the earliest synthetic speech sounds. This led to further work by Thomas Edison to record audio with <a href="Dictation_machines" class="mw-redirect" title="Dictation machines">dictation machines</a> and play it back in corporate settings. In the 1950s-1960s there were primitive attempts to build automated <a href="Speech_recognition" title="Speech recognition">speech recognition</a> systems by <a href="Bell_Labs" title="Bell Labs">Bell Labs</a>, <a href="IBM" title="IBM">IBM</a>, and others. However, it was not until the 1980s that <a href="Hidden_Markov_Models" class="mw-redirect" title="Hidden Markov Models">Hidden Markov Models</a> were used to recognize up to 1,000 words that speech recognition systems became relevant.
</p>
<table class="wikitable">
<tbody><tr>
<th>Date
</th>
<th>Event
</th></tr>
<tr>
<td>1784
</td>
<td><a href="Wolfgang_von_Kempelen" title="Wolfgang von Kempelen">Wolfgang von Kempelen</a> creates the Acoustic-Mechanical speech machine.
</td></tr>
<tr>
<td>1879
</td>
<td><a href="Thomas_Edison" title="Thomas Edison">Thomas Edison</a> invents the first <a href="Dictation_machine" title="Dictation machine">dictation machine</a>.
</td></tr>
<tr>
<td>1952
</td>
<td><a href="Bell_Labs" title="Bell Labs">Bell Labs</a> releases <a href="Audrey" title="Audrey">Audrey</a>, capable of recognizing spoken digits with 90% accuracy.
</td></tr>
<tr>
<td>1962
</td>
<td><a href="IBM_Shoebox" title="IBM Shoebox">IBM Shoebox</a> can recognize up to 16 words.
</td></tr>
<tr>
<td>1971
</td>
<td><a href="Harpy" title="Harpy">Harpy</a> is created, which can understand over 1,000 words.
</td></tr>
<tr>
<td>1986
</td>
<td>IBM Tangora uses <a href="Hidden_Markov_Models" class="mw-redirect" title="Hidden Markov Models">Hidden Markov Models</a> to predict phonemes in speech.
</td></tr>
<tr>
<td>2006
</td>
<td><a href="National_Security_Agency" title="National Security Agency">National Security Agency</a> begins research in hotword detection during normal conversations.
</td></tr>
<tr>
<td>2008
</td>
<td><a href="Google" title="Google">Google</a> launches a voice application, bring speech recognition to mobile devices.
</td></tr>
<tr>
<td>2011
</td>
<td><a href="Apple_Inc." title="Apple Inc.">Apple</a> releases Siri on iPhone
</td></tr>
<tr>
<td>2014
</td>
<td><a href="Amazon_(company)" title="Amazon (company)">Amazon</a> releases <a href="Amazon_Echo" title="Amazon Echo">Amazon Echo</a> to make voice computing relevant to the public at large.
</td></tr></tbody></table>
<p>Around 2011, <a href="Siri" title="Siri">Siri</a> emerged on Apple iPhones as the first voice assistant accessible to consumers. This innovation led to a dramatic shift to building voice-first computing architectures. <a href="PS4" class="mw-redirect" title="PS4">PS4</a> was released by Sony in North America in 2013 (70+ million devices), Amazon released the <a href="Amazon_Echo" title="Amazon Echo">Amazon Echo</a> in 2014 (30+ million devices), <a href="Microsoft" title="Microsoft">Microsoft</a> released Cortana (2015 - 400 million Windows 10 users), Google released <a href="Google_Assistant" title="Google Assistant">Google Assistant</a> (2016 - 2 billion active monthly users on Android phones), and <a href="Apple_Inc." title="Apple Inc.">Apple</a> released <a href="HomePod" title="HomePod">HomePod</a> (2018 - 500,000 devices sold and 1 billion devices active with iOS/Siri). These shifts, along with advancements in cloud infrastructure (e.g. <a href="Amazon_Web_Services" title="Amazon Web Services">Amazon Web Services</a>) and <a href="Codecs" class="mw-redirect" title="Codecs">codecs</a>, have solidified the voice computing field and made it widely relevant to the public at large.
</p>
<div class="mw-heading mw-heading2"><h2 id="Hardware">Hardware</h2></div>
<p>A <b>voice computer</b> is assembled hardware and software to process voice inputs.
</p><p>Note that voice computers do not necessarily need a screen, such as in the traditional <a href="Amazon_Echo" title="Amazon Echo">Amazon Echo</a>. In other embodiments, traditional <a href="Laptop_computers" class="mw-redirect" title="Laptop computers">laptop computers</a> or <a href="Mobile_phones" class="mw-redirect" title="Mobile phones">mobile phones</a> could be used as voice computers. Moreover, there has become increasingly more interfaces for voice computers with the advent of <a href="Internet_of_things" title="Internet of things">IoT</a>-enabled devices, such as within cars or televisions.
</p><p>As of September 2018, there are currently over 20,000 types of devices compatible with Amazon Alexa.<sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Software">Software</h2></div>
<p>Voice computing software can read/write, record, clean, encrypt/decrypt, playback, transcode, transcribe, compress, publish, featurize, model, and visualize voice files.
</p><p>Here are some popular software packages related to voice computing:
</p>
<table class="wikitable">
<tbody><tr>
<th>Package name
</th>
<th>Description
</th></tr>
<tr>
<th><a href="FFmpeg" title="FFmpeg">FFmpeg</a>
</th>
<td>for <a href="Transcoding" title="Transcoding">transcoding</a> audio files from one format to another (e.g. .WAV --> .MP3).<sup id="cite_ref-4" class="reference"><a href="#cite_note-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th><a href="Audacity_(audio_editor)" title="Audacity (audio editor)">Audacity</a>
</th>
<td>for recording and filtering audio.<sup id="cite_ref-5" class="reference"><a href="#cite_note-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th><a href="SoX" title="SoX">SoX</a>
</th>
<td>for manipulating audio files and removing environmental noise.<sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th><a href="Natural_Language_Toolkit" title="Natural Language Toolkit">Natural Language Toolkit</a>
</th>
<td>for featurizing transcripts with things like <a href="Parts_of_speech" class="mw-redirect" title="Parts of speech">parts of speech</a>.<sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th>LibROSA
</th>
<td>for visualizing audio file spectrograms and featurizing audio files.<sup id="cite_ref-8" class="reference"><a href="#cite_note-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th><a href="OpenSMILE" title="OpenSMILE">OpenSMILE</a>
</th>
<td>for featurizing audio files with things like mel-frequency cepstrum coefficients.<sup id="cite_ref-9" class="reference"><a href="#cite_note-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th><a href="CMU_Sphinx" title="CMU Sphinx">CMU Sphinx</a>
</th>
<td>for transcribing speech files into text.<sup id="cite_ref-10" class="reference"><a href="#cite_note-10"><span class="cite-bracket">[</span>10<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th>Pyttsx3
</th>
<td>for playing back audio files (text-to-speech).<sup id="cite_ref-11" class="reference"><a href="#cite_note-11"><span class="cite-bracket">[</span>11<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th>Pycryptodome
</th>
<td>for encrypting and decrypting audio files.<sup id="cite_ref-12" class="reference"><a href="#cite_note-12"><span class="cite-bracket">[</span>12<span class="cite-bracket">]</span></a></sup>
</td></tr>
<tr>
<th>AudioFlux
</th>
<td>for audio and music analysis, feature extraction.<sup id="cite_ref-13" class="reference"><a href="#cite_note-13"><span class="cite-bracket">[</span>13<span class="cite-bracket">]</span></a></sup>
</td></tr></tbody></table>
<div class="mw-heading mw-heading2"><h2 id="Applications">Applications</h2></div>
<p>Voice computing applications span many industries including voice assistants, healthcare, e-Commerce, finance, supply chain, agriculture, text-to-speech, security, marketing, customer support, recruiting, cloud computing, microphones, speakers, and podcasting. Voice technology is projected to grow at a CAGR of 19-25% by 2025, making it an attractive industry for startups and investors alike.<sup id="cite_ref-14" class="reference"><a href="#cite_note-14"><span class="cite-bracket">[</span>14<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Legal_considerations">Legal considerations</h2></div>
<p>In the United States, the states have varying <a href="Telephone_call_recording_laws" title="Telephone call recording laws">telephone call recording laws</a>. In some states, it is legal to record a conversation with the consent of only one party, in others the consent of all parties is required.
</p><p>Moreover, <a href="COPPA" class="mw-redirect" title="COPPA">COPPA</a> is a significant law to protect minors using the Internet. With an increasing number of minors interacting with voice computing devices (e.g. the Amazon Alexa), on October 23, 2017 the <a href="Federal_Trade_Commission" title="Federal Trade Commission">Federal Trade Commission</a> relaxed the COPAA rule so that children can issue voice searches and commands.<sup id="cite_ref-15" class="reference"><a href="#cite_note-15"><span class="cite-bracket">[</span>15<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-16" class="reference"><a href="#cite_note-16"><span class="cite-bracket">[</span>16<span class="cite-bracket">]</span></a></sup>
</p><p>Lastly, <a href="GDPR" class="mw-redirect" title="GDPR">GDPR</a> is a new European law that governs the <a href="Right_to_be_forgotten" title="Right to be forgotten">right to be forgotten</a> and many other clauses for EU citizens. GDPR also is clear that companies need to outline clear measures to obtain consent if audio recordings are made and define the purpose and scope as to how these recordings will be used, e.g., for training purposes. The bar for valid consent has been raised under the GDPR. Consents must be freely given, specific, informed, and unambiguous; tacit consent is no longer sufficient.<sup id="cite_ref-17" class="reference"><a href="#cite_note-17"><span class="cite-bracket">[</span>17<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Research_conferences">Research conferences</h2></div>
<p>There are many research conferences that relate to voice computing. Some of these include:
</p>
<ul><li><a href="International_Conference_on_Acoustics%2C_Speech%2C_and_Signal_Processing" title="International Conference on Acoustics, Speech, and Signal Processing">International Conference on Acoustics, Speech, and Signal Processing</a></li>
<li>Interspeech <sup id="cite_ref-18" class="reference"><a href="#cite_note-18"><span class="cite-bracket">[</span>18<span class="cite-bracket">]</span></a></sup></li>
<li>AVEC <sup id="cite_ref-19" class="reference"><a href="#cite_note-19"><span class="cite-bracket">[</span>19<span class="cite-bracket">]</span></a></sup></li>
<li>IEEE Int'l Conf. on Automatic Face and Gesture Recognition <sup id="cite_ref-20" class="reference"><a href="#cite_note-20"><span class="cite-bracket">[</span>20<span class="cite-bracket">]</span></a></sup></li>
<li>ACII2019 The 8th Int'l Conf. on Affective Computing and Intelligent Interaction <sup id="cite_ref-21" class="reference"><a href="#cite_note-21"><span class="cite-bracket">[</span>21<span class="cite-bracket">]</span></a></sup></li></ul>
<div class="mw-heading mw-heading2"><h2 id="Developer_community">Developer community</h2></div>
<p>Google Assistant has roughly 2,000 actions as of January 2018.<sup id="cite_ref-22" class="reference"><a href="#cite_note-22"><span class="cite-bracket">[</span>22<span class="cite-bracket">]</span></a></sup>
</p><p>There are over 50,000 Alexa skills worldwide as of September 2018.<sup id="cite_ref-23" class="reference"><a href="#cite_note-23"><span class="cite-bracket">[</span>23<span class="cite-bracket">]</span></a></sup>
</p><p>In June 2017, <a href="Google" title="Google">Google</a> released AudioSet,<sup id="cite_ref-24" class="reference"><a href="#cite_note-24"><span class="cite-bracket">[</span>24<span class="cite-bracket">]</span></a></sup> a large-scale collection of human-labeled 10-second sound clips drawn from YouTube videos. It contains 1,010,480 videos of human speech files, or 2,793.5 hours in total.<sup id="cite_ref-25" class="reference"><a href="#cite_note-25"><span class="cite-bracket">[</span>25<span class="cite-bracket">]</span></a></sup> It was released as part of the IEEE ICASSP 2017 Conference.<sup id="cite_ref-26" class="reference"><a href="#cite_note-26"><span class="cite-bracket">[</span>26<span class="cite-bracket">]</span></a></sup>
</p><p>In November 2017, <a href="Mozilla_Foundation" title="Mozilla Foundation">Mozilla Foundation</a> released the Common Voice Project, a collection of speech files to help contribute to the larger open source machine learning community.<sup id="cite_ref-27" class="reference"><a href="#cite_note-27"><span class="cite-bracket">[</span>27<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-28" class="reference"><a href="#cite_note-28"><span class="cite-bracket">[</span>28<span class="cite-bracket">]</span></a></sup> The voicebank is currently 12GB in size, with more than 500 hours of English-language voice data that have been collected from 112 countries since the project's inception in June 2017.<sup id="cite_ref-29" class="reference"><a href="#cite_note-29"><span class="cite-bracket">[</span>29<span class="cite-bracket">]</span></a></sup> This dataset has already resulted in creative projects like the DeepSpeech model, an open source transcription model.<sup id="cite_ref-30" class="reference"><a href="#cite_note-30"><span class="cite-bracket">[</span>30<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<ul><li><a href="Speech_recognition" title="Speech recognition">Speech recognition</a></li>
<li><a href="Natural_language_processing" title="Natural language processing">Natural language processing</a></li>
<li><a href="Voice_user_interface" title="Voice user interface">Voice user interface</a></li>
<li><a href="Audio_codec" title="Audio codec">Audio codec</a></li>
<li><a href="Ubiquitous_computing" title="Ubiquitous computing">Ubiquitous computing</a></li>
<li><a href="Hands-free_computing" title="Hands-free computing">Hands-free computing</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */
.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}
/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap mw-references-columns"><ol class="references">
<li id="cite_note-1"><span class="mw-cite-backlink"><b><a href="#cite_ref-1">^</a></b></span> <span class="reference-text">Schwoebel, J. (2018). An Introduction to Voice Computing in Python. Boston; Seattle, Atlanta: NeuroLex Laboratories. <a rel="nofollow" class="external free" href="https://neurolex.ai/voicebook">https://neurolex.ai/voicebook</a></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */
.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}
/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFBoyd2019" class="citation web cs1">Boyd, Clark (2019-08-30). <a rel="nofollow" class="external text" href="https://medium.com/swlh/the-past-present-and-future-of-speech-recognition-technology-cf13c179aaf">"Speech Recognition Technology: The Past, Present, and Future"</a>. <i>The Startup</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><b><a href="#cite_ref-3">^</a></b></span> <span class="reference-text"><cite id="CITEREFKinsella2018" class="citation web cs1">Kinsella, Bret (2018-09-02). <a rel="nofollow" class="external text" href="https://voicebot.ai/2018/09/02/amazon-alexa-now-has-50000-skills-worldwide-is-on-20000-devices-used-by-3500-brands/">"Amazon Alexa Now Has 50,000 Skills Worldwide, works with 20,000 Devices, Used by 3,500 Brands"</a>. <i>Voicebot.ai</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-4">^</a></b></span> <span class="reference-text">FFmpeg. <a rel="nofollow" class="external free" href="https://www.ffmpeg.org/">https://www.ffmpeg.org/</a></span>
</li>
<li id="cite_note-5"><span class="mw-cite-backlink"><b><a href="#cite_ref-5">^</a></b></span> <span class="reference-text">Audacity. <a rel="nofollow" class="external free" href="https://www.audacityteam.org/">https://www.audacityteam.org/</a></span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-6">^</a></b></span> <span class="reference-text">SoX. <a rel="nofollow" class="external free" href="https://sox.sourceforge.net/">https://sox.sourceforge.net/</a></span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><b><a href="#cite_ref-7">^</a></b></span> <span class="reference-text">NLTK. <a rel="nofollow" class="external free" href="https://www.nltk.org/">https://www.nltk.org/</a></span>
</li>
<li id="cite_note-8"><span class="mw-cite-backlink"><b><a href="#cite_ref-8">^</a></b></span> <span class="reference-text">LibROSA. <a rel="nofollow" class="external free" href="https://librosa.github.io/librosa/">https://librosa.github.io/librosa/</a></span>
</li>
<li id="cite_note-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-9">^</a></b></span> <span class="reference-text">OpenSMILE. <a rel="nofollow" class="external free" href="https://www.audeering.com/technology/opensmile/">https://www.audeering.com/technology/opensmile/</a></span>
</li>
<li id="cite_note-10"><span class="mw-cite-backlink"><b><a href="#cite_ref-10">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://github.com/cmusphinx/pocketsphinx">"PocketSphinx is a lightweight speech recognition engine, specifically tuned for handheld and mobile devices, though it works equally well on the desktop: Cmusphinx/Pocketsphinx"</a>. <i><a href="GitHub" title="GitHub">GitHub</a></i>. 29 March 2020.</cite></span>
</li>
<li id="cite_note-11"><span class="mw-cite-backlink"><b><a href="#cite_ref-11">^</a></b></span> <span class="reference-text">Pyttsx3. <a rel="nofollow" class="external free" href="https://github.com/nateshmbhat/pyttsx3">https://github.com/nateshmbhat/pyttsx3</a></span>
</li>
<li id="cite_note-12"><span class="mw-cite-backlink"><b><a href="#cite_ref-12">^</a></b></span> <span class="reference-text">Pycryptodome. <a rel="nofollow" class="external free" href="https://pycryptodome.readthedocs.io/en/latest/">https://pycryptodome.readthedocs.io/en/latest/</a></span>
</li>
<li id="cite_note-13"><span class="mw-cite-backlink"><b><a href="#cite_ref-13">^</a></b></span> <span class="reference-text">AudioFlux. <a rel="nofollow" class="external free" href="https://github.com/libAudioFlux/audioFlux/">https://github.com/libAudioFlux/audioFlux/</a></span>
</li>
<li id="cite_note-14"><span class="mw-cite-backlink"><b><a href="#cite_ref-14">^</a></b></span> <span class="reference-text"><cite class="citation news cs1"><a rel="nofollow" class="external text" href="http://web.archive.org/web/20240119171935/https://www.businesswire.com/news/home/20180417006122/en/Global-Speech-Voice-Recognition-Market-2018-Forecast">"Global Speech and Voice Recognition Market 2018 Forecast to 2025 - CAGR Expected to Grow at 25.7% - ResearchAndMarkets.com"</a>. Archived from <a rel="nofollow" class="external text" href="https://www.businesswire.com/news/home/20180417006122/en/Global-Speech-Voice-Recognition-Market-2018-Forecast">the original</a> on 2024-01-19<span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-15"><span class="mw-cite-backlink"><b><a href="#cite_ref-15">^</a></b></span> <span class="reference-text"><cite id="CITEREFColdewey2017" class="citation web cs1">Coldewey, Devin (2017-10-24). <a rel="nofollow" class="external text" href="https://techcrunch.com/2017/10/24/ftc-relaxes-coppa-rule-so-kids-can-issue-voice-searches-and-commands/">"FTC relaxes COPPA rule so kids can issue voice searches and commands"</a>. <i>TechCrunch</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-16"><span class="mw-cite-backlink"><b><a href="#cite_ref-16">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.federalregister.gov/documents/2017/12/08/2017-26509/enforcement-policy-statement-regarding-the-applicability-of-the-coppa-rule-to-the-collection-and-use">"Federal Register :: Request Access"</a>. 8 December 2017.</cite></span>
</li>
<li id="cite_note-17"><span class="mw-cite-backlink"><b><a href="#cite_ref-17">^</a></b></span> <span class="reference-text">IAPP. <a rel="nofollow" class="external free" href="https://iapp.org/news/a/how-do-the-rules-on-audio-recording-change-under-the-gdpr/">https://iapp.org/news/a/how-do-the-rules-on-audio-recording-change-under-the-gdpr/</a></span>
</li>
<li id="cite_note-18"><span class="mw-cite-backlink"><b><a href="#cite_ref-18">^</a></b></span> <span class="reference-text">Interspeech 2018. <a rel="nofollow" class="external free" href="http://interspeech2018.org/">http://interspeech2018.org/</a></span>
</li>
<li id="cite_note-19"><span class="mw-cite-backlink"><b><a href="#cite_ref-19">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://www.eventyco.com/event/14th-international-symposium-on-advanced-vehicle-control">"14th International Symposium on Advanced Vehicle Control - Speakers, Sessions, Agenda"</a>. <i>www.eventyco.com</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-20"><span class="mw-cite-backlink"><b><a href="#cite_ref-20">^</a></b></span> <span class="reference-text">2018 FG. <a rel="nofollow" class="external free" href="https://fg2018.cse.sc.edu/">https://fg2018.cse.sc.edu/</a></span>
</li>
<li id="cite_note-21"><span class="mw-cite-backlink"><b><a href="#cite_ref-21">^</a></b></span> <span class="reference-text">ASCII 2019. <a rel="nofollow" class="external free" href="http://acii-conf.org/2019/">http://acii-conf.org/2019/</a></span>
</li>
<li id="cite_note-22"><span class="mw-cite-backlink"><b><a href="#cite_ref-22">^</a></b></span> <span class="reference-text"><cite id="CITEREFMutchler2018" class="citation web cs1">Mutchler, Ava (2018-01-24). <a rel="nofollow" class="external text" href="https://voicebot.ai/2018/01/24/google-assistant-app-total-reaches-nearly-2400-thats-not-real-number-really-1719/">"Google Assistant App Total Reaches Nearly 2400. But That's Not the Real Number. It's really 1719"</a>. <i>Voicebot.ai</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-23"><span class="mw-cite-backlink"><b><a href="#cite_ref-23">^</a></b></span> <span class="reference-text"><cite id="CITEREFKinsella2018" class="citation web cs1">Kinsella, Bret (2018-09-02). <a rel="nofollow" class="external text" href="https://voicebot.ai/2018/09/02/amazon-alexa-now-has-50000-skills-worldwide-is-on-20000-devices-used-by-3500-brands/">"Amazon Alexa Now Has 50,000 Skills Worldwide, works with 20,000 Devices, Used by 3,500 Brands"</a>. <i>Voicebot.ai</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-24"><span class="mw-cite-backlink"><b><a href="#cite_ref-24">^</a></b></span> <span class="reference-text">Google AudioSet. <a rel="nofollow" class="external free" href="https://research.google.com/audioset/">https://research.google.com/audioset/</a></span>
</li>
<li id="cite_note-25"><span class="mw-cite-backlink"><b><a href="#cite_ref-25">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://research.google.com/audioset/dataset/speech.html">"AudioSet"</a>. <i>research.google.com</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-26"><span class="mw-cite-backlink"><b><a href="#cite_ref-26">^</a></b></span> <span class="reference-text">Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, & Ritter, M. (2017, March). Audio set: An ontology and human-labeled dataset for audio events. In Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on (pp. 776-780). IEEE.</span>
</li>
<li id="cite_note-27"><span class="mw-cite-backlink"><b><a href="#cite_ref-27">^</a></b></span> <span class="reference-text">Common Voice Project. <a rel="nofollow" class="external free" href="https://voice.mozilla.org/">https://voice.mozilla.org/</a></span>
</li>
<li id="cite_note-28"><span class="mw-cite-backlink"><b><a href="#cite_ref-28">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="https://blog.mozilla.org/en/mozilla/announcing-the-initial-release-of-mozillas-open-source-speech-recognition-model-and-voice-dataset/">"Announcing the Initial Release of Mozilla's Open Source Speech Recognition Model and Voice Dataset | The Mozilla Blog"</a>. <i>blog.mozilla.org</i><span class="reference-accessdate">. Retrieved <span class="nowrap">2025-01-10</span></span>.</cite></span>
</li>
<li id="cite_note-29"><span class="mw-cite-backlink"><b><a href="#cite_ref-29">^</a></b></span> <span class="reference-text">Mozilla's large repository of voice data will shape the future of machine learning. <a rel="nofollow" class="external free" href="https://opensource.com/article/18/4/common-voice">https://opensource.com/article/18/4/common-voice</a></span>
</li>
<li id="cite_note-30"><span class="mw-cite-backlink"><b><a href="#cite_ref-30">^</a></b></span> <span class="reference-text">DeepSpeech. <a rel="nofollow" class="external free" href="https://github.com/mozilla/DeepSpeech">https://github.com/mozilla/DeepSpeech</a></span>
</li>
</ol></div></div></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-01-10" href="https://en.wikipedia.org/wiki/?title=Voice_computing&oldid=1268559446">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
</body></html>